Papers with Augmented LibriSpeech
CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets involve language pairs with English as source language, are low resource or lack labeled data. |
| Approach: | They propose a multilingual speech-to-text translation corpus from 11 languages into English . they provide empirical evidence of the quality of the data and provide initial benchmarks . |
| Outcome: | The proposed model is the first end-to-end multilingual model for spoken language translation. |